Papers with fine-grained evaluation metrics
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics (2026.acl-long)
Copied to clipboard
| Challenge: | Web applications (web apps) are a key arena for large language models to demonstrate their code generation capabilities and commercial potential. |
| Approach: | a new benchmark for large language models (LLMs) is designed to provide real-world user requirements and generalizable evaluation metrics. |
| Outcome: | a new benchmark for large language models (LLMs) provides a real-world, generalizable, and interpretable evaluation score . the benchmark measures user requirements, expression styles and human-preference-aligned weights . a web application can be used to demonstrate its commercial potential, authors say . |